Papers with phrase-to-phrase supervision
A Visually-Grounded Parallel Corpus with Phrase-to-Region Linking (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing multimodal corpora lack the ability to be used in multilingual or non-English scenarios. |
| Approach: | They extend a Flickr30k Entities image-caption dataset with Japanese translations to provide a multilingual corpus. |
| Outcome: | The proposed dataset is the first multilingual image-caption dataset with Japanese translations. |